npj Systems Biology and Applications
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match npj Systems Biology and Applications's content profile, based on 125 papers previously published here. The average preprint has a 0.09% match score for this journal, so anything above that is already an above-average fit.
BV, H.; Adigwe, S.; Jolly, M. K.; Gedeon, T.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWCell fate decisions are driven by gene regulatory networks (GRNs). While the mutually inhibitory toggle switch effectively models binary fate decisions, fully connected inhibitory networks with more than two nodes fail to capture multi-fate decisions due to the low prevalence of "single high states", where only a single master regulator is highly expressed. The goal of this study is to find network structures that support all single high states. We find that the only network that attains the highest possible prevalence of all single high states within the set of monotone Boolean (MB) models is completely disconnected. Since biological networks typically require connectivity, we investigate network structures that support equipotency, where all single high states have equal prevalence within MB models. Finally, we characterize the networks that support multistability between all single high states, finding that it is possible only in networks in which each node either has self-activations or is inhibited by every other network node. Our findings provide a theoretical framework for understanding the network design principles that can support simultaneous differentiation into multiple distinct cell types.
Ozkurt, C.
Show abstract
BackgroundMicroglia drive neuroinflammation in Alzheimers disease (AD), yet no approved therapy targets this compartment. Human genome-wide association studies consistently implicate innate immune loci in AD risk, establishing microglial transcriptional programs as therapeutically relevant but pharmacologically underexploited targets. ObjectiveWe sought to identify transcription factors (TFs) governing microglial state transitions computationally and to nominate structurally tractable drug repurposing candidates. MethodsWe applied trajectory inference (PAGA), pseudobulk DESeq2, pySCENIC gene regulatory network (GRN) inference, CellChat, and virtual screening of 1,962 approved compounds to 236,002 microglial nuclei from 84 donors (SEA-AD atlas). ResultsIKZF1 was the sole target TF retained under cisTarget v10 motif constraints, with peak regulon activity in LateAD-DAM (pseudotime {rho} = +0.309) and replication in an independent bulk cohort (GSE95587; adjusted P value =.004). CellChat identified SLIT2[->]ROBO2 from multiple neuron subtypes (predominantly inhibitory interneurons) as the top predicted pathway to microglia. Tafamidis ([->]IRF8) and diflunisal ([->]PPARG) were top virtual screening hits; all evaluated compounds failed the pre-specified selectivity threshold. ConclusionsIKZF1 is prioritised as a candidate late-disease microglial TF, supported by six convergent evidence dimensions including independent bulk replication. Tafamidis and diflunisal are low-confidence repurposing hypotheses requiring experimental validation.
Adams, S.; Phelan, L.; Lewis, T.; Behm, J.; Law, A.; Shi, X.; Li, G. F.; Li, J.
Show abstract
Bipolar androgen therapy (BAT) exploits the paradoxical vulnerability of castration-resistant prostate cancer (CRPC) cells to rapid cycling between castrate and supraphysiologic androgen concentrations, but clinical BAT uses testosterone, which can also activate wild-type androgen receptor (AR) in androgen-responsive tissues, causing systemic side effects. 5{beta}-dihydrotestosterone (5{beta}-DHT) is a naturally occurring testosterone metabolite generally considered androgenically inactive because it binds wild-type AR weakly, yet its activity against clinically relevant AR mutants has not been systematically evaluated. Here, we tested whether 5{beta}-DHT and related 5{beta}-reduced testosterone metabolites activate AR signaling and growth programs in prostate cancer models that carry AR mutations. In C4-2 cells, 5{beta}-DHT and 3{beta}-etiocholanediol (3{beta}-ecdiol) increased canonical AR target genes, including KLK3 and TMPRSS2, with weaker activity than testosterone, whereas other 5{beta} metabolites showed limited activity. In androgen-responsive LNCaP and C4-2 models, 5{beta}-DHT and 3{beta}-ecdiol promoted cell growth under androgen-depleted conditions, and this effect was suppressed by enzalutamide, supporting AR dependence. RNA-seq confirmed that 5{beta}-DHT and 3{beta}-ecdiol induced androgen-response gene sets substantially overlapping with testosterone, albeit at lower transcriptional magnitude. Further, we found that 5{beta}-DHT, but not 3{beta}-ecdiol, suppresses cell proliferation of LNCaP, C4-2, and PC-3 cells stably expressing the clinically relevant AR gain-of-function mutants W742C and H875Y through activating AR-induced senescence-like features after high-dose exposure, consistent with the therapeutic logic of BAT. These findings identify 5{beta}-DHT as an overlooked mutant-AR agonist capable of BAT-like tumor suppression and propose it as a testosterone surrogate in BAT with potentially reduced systemic androgenic side effects. HighlightsO_LI5{beta}-DHT and 3{beta}-ecdiol promote AR-dependent prostate cancer cell growth C_LIO_LIBoth are weaker AR agonists than testosterone by RNA-seq and qPCR C_LIO_LISupraphysiologic 5{beta}-DHT suppresses growth via AR-mediated senescence C_LIO_LIGrowth suppression extends to AR mutants W742C and H875Y C_LIO_LI5{beta}-DHT may be a lower-androgenicity testosterone surrogate for BAT C_LI
Kumar Reddy, K.; Hahn, W.; Winter, S.; Roellig, C.; Mueller-Tidow, C.; Serve, H.; Baldus, C. D.; Fransecky, L.; Schliemann, C.; Burchert, A.; Schaefer-Eckart, K.; Kaufmann, M.; Schetelig, J.; Bornhaeuser, M.; Middeke, J. M.; Eckardt, J.-N.
Show abstract
Rising costs, slow accrual and molecular substratification of cancers necessitate novel clinical trial designs. We demonstrate that artificial intelligence-generated synthetic patients can replace real controls to reproduce results of the SORAML trial. Using external multimodal data from 1,377 acute myeloid leukemia (AML) patients from previous trials and a real-world registry, we fine-tuned a tabular foundation model to generate synthetic patients, reproducing clinical and genetic features and outcome associations. Synthetic patients were then matched to the original SORAML intervention group using Cox risk scores, replacing the original control and reproducing the original trial result with near-identical median event-free survival (EFS) and treatment effect (original hazard ratio [HR] 0.64, 95%-confidence interval [CI] 0.47-0.87, p=0.004; with synthetic control HR 0.66, 95%-CI 0.48-0.90, p=0.009). Our findings demonstrate that AI-generated synthetic patients can serve as statistically rigorous controls supporting novel trial designs.
Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.
Show abstract
Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.
Pepe, M.; Hesami, M.; Jones, M.
Show abstract
Applications of tissue culture are critical for Cannabis sativa L. (cannabis), supporting clonal propagation, germplasm preservation, pathogen elimination, among other biotechnological applications. However, extensive genetic diversity associated with cannabis results in highly variable responses to in vitro conditioning, and no consensus basal media formulation exists to support reproducible micropropagation across genotypes. To address these limitations, a hybridized ensemble-NSGA-II approach was employed for concurrent optimization of individual media components to create a species specific, cultivar inclusive basal salt formulation for cannabis micropropagation. The resulting PHJ media represents a unique formulation that overcomes recalcitrance across a wide array of cannabis cultivars, facilitating improved growth and uniformity for the nine cultivars used in its development and validation. These results remain consistent from explant initiation through multiple rounds of subculture. The ability of PHJ to overcome genotypic recalcitrance is telling of its potential applicability with an array of plant species beyond cannabis. Additionally, robust performance both with and without plant growth regulators underscores the plausible use of PHJ for diverse applications beyond standard micropropagation. Ultimately, this cultivar-inclusive basal medium demonstrates utility for both scientific research and industrial-scale operations.
Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.
Show abstract
Background Unsupervised machine learning has become a cornerstone of computational phenotyping across clinical medicine, genomics, imaging, and multi-omics research. However, phenotype discovery relies on a sequence of analytical decisions - including missing-data handling, preprocessing, dimensionality reduction, clustering methodology, and stochastic initialization - that are rarely evaluated collectively. Although clustering stability has been extensively investigated, the robustness of complete analytical workflows remains largely unexplored. Results We developed an Analytical Perturbation Framework that systematically quantifies the robustness of phenotype discovery by perturbing complete unsupervised learning workflows rather than individual clustering algorithms. Using a real-world cohort of 1,286 women with polycystic ovary syndrome (PCOS), we generated 116 valid analytical pipelines comprising alternative preprocessing strategies, missing-data handling methods, dimensionality reduction approaches, clustering algorithms, and random initializations. Agreement between independently generated phenotype solutions was consistently low (median Adjusted Rand Index = 0.079), indicating substantial sensitivity of phenotype discovery to routine analytical decisions. Variance decomposition identified preprocessing as the largest contributor to phenotype instability (22.8%), followed by clustering methodology (14.6%), whereas stochastic initialization explained only 3.1% of the observed variability. At the patient level, most individuals exhibited reproducible phenotype assignments (median Patient Robustness Score = 0.719), although a substantial subgroup showed markedly lower assignment stability. Feature perturbation analyses identified follicle-stimulating hormone, anti-thyroglobulin antibodies, anti-thyroid peroxidase antibodies, total testosterone, luteinizing hormone, and androstenedione as the strongest contributors to computational robustness, rather than biological importance. Finally, phenotype solutions demonstrating greater computational robustness also exhibited greater biological coherence during independent validation.
Bryan, C. B.; Kilic, F.; Garcia, I.; Ly, A.; Ly, A.; Muhammad, A.; Kwok, H. Y.; Miranda, V.; Bashar, A.; Polagoni, A.; Bacchus, Z.; Yang, K.; Klein, E. A.; Corbett, B. F.
Show abstract
Stress-related psychiatric disorders and inflammatory bowel diseases share high co-morbidity and contribute to the symptom severity of one another. In mice, ten days of Chronic Social Defeat Stress (CSDS) is sufficient to reduce gut microbiome diversity and the relative abundance of Firmicutes, which are hallmarks of inflammatory bowel diseases. However, mechanisms by which stress causes gut microbiome dysbiosis are largely unknown. Here, we demonstrate that pharmacologically inhibiting {beta}-adrenergic receptors (ARs), which are activated by (nor)adrenaline during stress, mitigates gut dysbiosis otherwise caused by CSDS. Compared to vehicle-treated mice following CSDS, propranolol-treated mice displayed a modest increase in sociability, increased alpha diversity, and increased abundance of anaerobic commensal Clostridia. Abundance of short-chain fatty acid-producing anaerobic Firmicutes abundance correlated with sociability following CSDS across all treatments. Pharmacologically blocking -ARs during stress increased subsequent sociability, but had little effect on gut microbiome composition. Together, our findings support the hypothesis that {beta}-AR activation contributes to stress-induced changes of the gut microbiome. One Sentence SummaryPharmacologically inhibiting beta-adrenergic receptors during chronic stress mitigates reductions in anaerobic, short-chain fatty acid-producing bacteria in the gut.
Cheng, C.
Show abstract
Genome-scale Perturb-seq screens prioritize candidate targets by the strength of a perturbations transcriptional effect. Effect strength does not answer a prior measurement question: is the readout dependable? A large effect estimated from a single guide, a single donor, or a pseudobulk of few cells need not survive replication, and for target prioritization each false lead costs a validation experiment. We treat each perturbation effect as a measurement in a crossed Target x Guide x Donor x Condition design and apply generalizability theory (Brennan, 2001; Cronbach et al., 1972) to separate the dependable part of an effect from facet-specific idiosyncrasy. Guides and donors enter as random facets; condition enters as a fixed facet and is analyzed within its levels. For each target we report a dependability profile over the facets and a joint generalizability coefficient over the two random facets, and we re-rank targets by effect magnitude weighted by that coefficient. On the released screen (Zhu et al., 2025), removing the measurement-error floor estimated from the non-targeting controls raises the number of genes with a dependable target-signal share above .10 from 40 to 7,674. Analyzed within activation states, dependability recovers the T-cell-receptor signaling module as reliably measurable only in activated cells, without recourse to gene annotation. A design study indicates that reliability is limited by the number of guides rather than the number of donors, so a future screen should add guides. Every methodological decision was recorded and adversarially reviewed, and all results regenerate from the released summary statistics.
Bueso, F. G.; Wardle, R.; Manescu, P.; Spear, J.; Ray, S.; Peters, M.
Show abstract
Offline reinforcement learning (RL) has emerged as a promising framework for clinical decision support in sepsis, yet most existing studies focus exclusively on adult populations, leaving pediatric care largely unexplored despite important physiological and treatment differences. In this work, we develop offline RL policies for pediatric sepsis management in the Pediatric Intensive Care Unit (PICU) using a retrospective cohort of 2,229 episodes from Great Ormond Street Hospital (GOSH), formalized as finite horizon Markov Decision Process (MDP) with joint intravenous fluid and vasopressor actions. To better capture pediatric organ dysfunction dynamics, we incorporate Phoenix 8, a recently proposed pediatric sepsis severity score, as an intermediate reward shaping signal in addition to terminal 90 day mortality. We systematically vary the time step size (4, 8, and 12 hours) and reward structure (terminal 90 day mortality, with and without Phoenix 8 based intermediate shaping), and compare Double Deep Q Networks (DDQN), Conservative Q Learning (CQL), and a behavior cloning (BC) model of clinician practice. CQL consistently exhibits stable learning dynamics and favorable Fitted Q Evaluation estimates, while DDQN is prone to overestimation and instability, particularly at finer temporal resolutions and with dense rewards. CQL policies achieve high action-level agreement with historical clinician decisions for both fluids and vasopressors and reproduce clinically plausible escalation patterns across sepsis severity strata, whereas DDQN policies diverge more frequently toward implausible dosing. Temporal aggregation emerges as a key regularizer: moving from 4 hour to 8 hour bins shortens horizons, smooths reward noise, and improves stability without erasing clinically meaningful dynamics, with 8 hour binning providing the best trade off between policy performance and granularity. Our findings highlight time step size as a core design choice in offline RL for healthcare and provide empirical evidence that alternatives beyond the conventional 4 hour setup can enhance stability and safety while preserving clinical interpretability.
Jabre, J. F.
Show abstract
The aim of this work is to validate patient-specific EEG baseline establishment using the e-norms method as a screening and retrospective-review tool for seizure detection in pediatric epilepsy. The method was applied to 247 seizure-free EEG recordings (263.92 hours) from 10 patients in the CHB-MIT Scalp EEG Database (ages 3-18). A composite stability metric combining first-derivative dynamics, spectral entropy, variance, and line length was computed per 2-second epoch across 23 channels. Patient-specific detection thresholds were derived from each patient's seizure-free baseline using a weighted statistical procedure. Performance was validated against 72 expert-annotated seizures (2,705 epochs) across 62 seizure files, with durations spanning 6 to 264 seconds (44-fold range). The results show that detection achieved 94.4% event-level sensitivity (68 of 72 seizures; 95% CI 86.6-97.8%) and 81.5% epoch-level sensitivity (2,204 of 2,705 epochs; 95% CI 80.0-82.9%). Eight of ten patients achieved 100% event-level sensitivity with epoch-level sensitivity ranging from 58.7% to 100.0%. Two patients showed partial event-level failures (CHB-15: 17 of 20; CHB-18: 5 of 6), with the four missed events attributable to two characterizable failure modes. Patient-specific thresholds ranged from 4.06 to 4.81 (mean 4.51 +/- 0.25); threshold variation did not correlate reliably with age or sex, confirming that no universal threshold could achieve comparable performance. Detection margins ranged from 0.88 to 1.24 times. Patient-specific e-norms achieves 94.4% event-level sensitivity for pediatric EEG seizure detection without requiring labeled seizure training data, exceeding published human expert inter-rater agreement (50-76%) and recent automated approaches in adult cohorts using behind-the-ear EEG and wearable ECG. Two characterizable failure modes account for the four missed events and inform appropriate clinical use. As a high-sensitivity screening tool complementary to real-time alarm systems, the method is ready for adult validation, prospective deployment, and head-to-head benchmarking.
Dangjarean, H.; Murata, Y.; Kobayashi, Y.; Neyrot, S.; Ogata, T.; Fujita, Y.
Show abstract
Plant-associated bacteria can improve plant performance under abiotic stress, but beneficial functions in plant microbiomes may depend on defined combinations of microorganisms rather than individual isolates alone. Here, we developed a cube-based screening strategy to identify functional synthetic microbial communities (SynComs) from 135 quinoa-associated bacterial isolates while preserving combinatorial diversity and traceability of isolate-level contributions. The isolates were divided into five 27-isolate sets, each arranged as a 3 x 3 x 3 cube in which each 3 x 3 layer was defined as a 9-isolate SynCom, generating 45 SynComs in total. Screening under 100 mM NaCl identified SynCom DY1 (SCDY1) as a candidate salt stress-mitigating consortium. SCDY1 consisted of nine taxonomically diverse isolates and exhibited a multifunctional profile, including siderophore production, phosphate solubilization, carboxymethyl cellulose degradation, indole compound production, and growth under saline conditions. In Arabidopsis thaliana, SCDY1 promoted primary root elongation and biomass accumulation in a salinity-dependent manner, with the clearest effect under 120 mM NaCl, and at least a subset of constituent bacteria was recoverable from inoculated seedlings. RNA sequencing and targeted RT-qPCR indicated that SCDY1 modulated host gene expression under moderate salinity stress, with responsive genes associated with oxidative stress, water- and oxygen-related processes, phenylpropanoid biosynthesis, glutathione metabolism, and root epidermis-related processes. Root hair phenotyping further showed that SCDY1 enhanced root hair-related traits and shifted visible root hair formation closer to the root apex. These findings identify a quinoa-derived SynCom that improves plant performance under salinity stress and provide a practical, traceable framework for discovering beneficial microbial consortia from plant-associated bacterial collections. Scope statementThis manuscript fits the Research Topic "Harnessing Plant Microbiomes for Climate Resilience: From Ecological Insight to Synthetic Community Design" in Frontiers in Plant Science because it presents a traceable strategy for discovering functional synthetic microbial communities from a stress-adapted plant-associated bacterial collection. We developed a cube-based screening strategy using 135 quinoa-associated bacterial isolates and identified a nine-isolate synthetic microbial community, SCDY1, that promotes Arabidopsis growth under moderate salinity stress. The study integrates microbiological screening, characterization of plant growth-promoting traits, bacterial re-isolation, plant growth phenotyping, RNA-seq, RT-qPCR, and root hair phenotyping. These analyses link SCDY1 treatment to salinity-dependent growth promotion, recoverable bacterial members, stress- and redox-associated transcriptional changes, phenylpropanoid-related responses, and modulation of root epidermal phenotypes. By connecting a defined SynCom with host transcriptional and root epidermal responses, this work advances understanding of beneficial plant-microbe interactions under salt stress. The cube-based design also provides a practical and traceable framework for discovering functional SynComs from large plant-associated bacterial collections, which should be of interest to researchers studying plant symbiosis, microbiome engineering, abiotic stress tolerance, and sustainable crop improvement.
Espero, M.
Show abstract
Background & Methods: The multifaceted physical nature of heritable cognitive impairment in dementia presents significant challenges for traditional linear frameworks attempting to model synergistic risk. While various loci are identified as contributing to neurocognitive disparities, the emergent phenotypic expression and associated predictive value relative to standard clinical baselines require further investigation. To facilitate dimensional reduction of complex genetic data into identifiable phenotypes, Generalized Low Rank Modeling (GLRM) and K-means clustering are applied to participant data from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The utility of these derived archetypes and clusters is assessed, stratifying variance for Mini-Mental State Examination (MMSE) performance. Utilizing generalized additive modeling (GAM) and partial eta squared (p2) effect size, the derived genetic features are compared with other predictors including age, educational attainment, gender, and raw, genetic variant carriage dimensions. Results & Conclusion: In accordance with the hypothesized empirical regularity, age and education persist as primary predictors of MMSE performance. The unsupervised machine learning pipeline successfully identified a composite genetic cluster that emerged as an influential predictor in terms of relative magnitude (p2). Centroid analysis of the GLRM subspace indicated that a particular sub-population (Cluster 2) - defined by a substantial weighting on the EPHA1 target - demonstrated a statistically significant association with MMSE scores, relative to cluster 3. These results suggest that data-driven genetic feature engineering provides an interpretable basis for inference regarding variance in global cognition. By discovering multivariate genetic architecture, this modeling approach captures complexity often missed by individual clinical variable modeling. Such findings implicate the utility of interpretable machine learning for translational dementia research and predictive clinical stratification.
Vasconcelos-Blomberg, P.; Felix China, J.; Syeda, B. R.; Fladvad, M.; Lagerlund, O.; Gattepaille, L. M.; Fusaroli, M.
Show abstract
Introduction: Conventional substance-level disproportionality analysis may miss safety patterns specific to a dose form, route, or intended site. More granular analyses are hindered by incomplete, inconsistent reporting of product information. The Pharmaceutical Product Identifier (PhPID), representing products by substance, strength, and dose form, may support more granular analyses. Objective: To explore the use of PhPID-like dose form information for site-specific disproportionality analysis in dexamethasone. Methods: We evaluated VigiBase reports (January 1, 2001 - December 31, 2024) for completeness of dose form and route data. We standardized dexamethasone entries to PhPID Level 3 standards, representing substance and administrable dose form. Through disproportionality analysis (Information Component, IC) we compared substance-level and site-specific results. Results: Among 56.4 million suspected/interacting drugs, dose form was reported in 47.7%, route in 69.4%. Among 109,248 dexamethasone entries, 703 dose form and 80 route variations were mapped to 53 and 44 standard codes respectively; about half could be mapped unambiguously. Site-specific analyses revealed biologically plausible patterns not apparent in substance-level analyses. Ocular use showed higher ICs for glaucoma and cataract, while systemic use showed higher IC for psychiatric and endocrine events (e.g., depression, agitation, Cushing's syndrome). IC time-trends suggested that some signals (e.g., cataract with Ocular use) could emerge earlier in site-specific analyses. Conclusion: More granular product information, aligned with PhPID, may improve signal detection and characterization of site-specific safety issues. These findings support granular identifiers in pharmacovigilance while highlighting the need for better capture and standardization of dose form and route of administration data.
Korutla, R.; Amal, S.
Show abstract
The Cancer Genome Atlas (TCGA) holds clinical data for over 11,000 patients across 33 cancer types, but access is hard because of complex file structures, heterogeneous formats, and the need for programming. We present an agentic system for natural language querying and statistical analysis of TCGA clinical data. The system uses a large language model as an autonomous ReAct agent that selects from eight computational tools, including data extraction, descriptive statistics, Kaplan-Meier survival analysis with log-rank tests, hypothesis testing, and verification against the curated TCGA Pan-Cancer Clinical Data Resource (CDR). The agent reasons about intermediate results, adapts its approach, and returns clinically contextualized responses with source attribution and auditable traces. We introduce TCGA-Agent-Bench, 440 queries across five difficulty tiers with ground truth from the independently curated TCGA-CDR, evaluated with dual metrics of numerical accuracy and clinical completeness. The system achieves 93.4% overall accuracy (100% single-patient lookups, 99.1% cohort statistics, 92.8% comparative analyses), outperforming a fixed rule-based pipeline (87.1%), a single-pass LLM (81.8%), and retrieval-augmented generation (66.9% on a subset). Most of the benchmark is answerable from the CDR alone, so we locate the extraction layer's value in fields the CDR lacks (drug treatments, TNM components, biomarkers, biospecimen metadata): on 26 queries targeting these, the full system answers 100% versus 3.8% for CDR-only. Ablations show the reasoning loop is most impactful (+9.1% accuracy, +22.0 completeness points). A tool-based agentic architecture enables accurate, auditable analysis of clinical repositories, with value driven by tool design and recovered fields rather than model scale.
Narisu, N.; Li, H. X.; Rathbun, C. J. M.; Varshney, A.; Swift, A. J.; Yan, T.; Sinha, N.; Currin, K. W.; Xue, D.; Robertson, C. C.; Taylor, D. L.; Taylor, H. J.; Beck, A.; Lee, B. N.; Wang, L.; Broadaway, K. A.; Wilson, E. P.; Stringham, H.; Saramies, J.; Lakka, T. A.; Spracklen, C. N.; Scott, L. J.; Stitzel, M. L.; Tuomilehto, J.; Laakso, M.; Koistinen, H. A.; Boehnke, M.; Arda, H. E.; Chen, S.; Biesecker, L. G.; Bonnycastle, L. L.; Erdos, M. R.; Mohlke, K. L.; Parker, S. C. J.; Collins, F. S.
Show abstract
Genome-wide association studies (GWAS) have identified >1,200 signals associated with type 2 diabetes (T2D), yet identifying functional variants remains challenging because the majority of them lie in noncoding regions of the genome and are in areas of high linkage disequilibrium (LD). While chromatin accessibility QTL (caQTL) and expression QTL (eQTL) analyses are useful for nominating regulatory mechanisms underlying GWAS signals, limitations still exist in pinpointing functional variants within regions of high LD. A complementary approach that has been less frequently applied is to focus on the allele-specific effect on chromatin accessibility at heterozygous single-nucleotide polymorphisms (SNPs), hereafter referred to as allelic imbalance. We analyzed the allelic imbalance of reads generated from an assay for transposase-accessible chromatin with sequencing (ATAC-seq) across genotyped samples from 490 donors in T2D-relevant tissues: skeletal muscle, liver, pancreatic islets, adipose tissue, and relevant cell types. We identified 119,949 allelically imbalanced SNPs (FDR<0.05) across the genome. The allelic imbalance was often most prominent in one tissue and showed an enrichment overlapping with tissue-specific transcription factor (TF) binding footprints. Focusing on the 8,581 SNPs in previously published 99% credible sets from 338 T2D GWAS signals, we identified 256 imbalanced SNPs across 123 (36.4% of) signals, each showing allelic imbalance in at least one tissue or cell type. Of these, 71 signals contained only a single imbalanced SNP, representing excellent candidate causative variants. As a proof-of-concept, we showed that 23 of the 256 imbalanced SNPs were supported by allelic assays from previous studies. Further, we experimentally validated two imbalanced SNPs as likely functional variants: rs34584161 among a seven-SNP T2D credible set at the RNF6 signal in islets and rs849134 among a 13-SNP credible set at the JAZF1 signal in liver. This study demonstrates the power of integrating ATAC-seq allelic imbalance (ASAI) with GWAS statistical fine-mapping to identify candidate functional regulatory variants from among tightly linked GWAS variants in disease-relevant tissues. While applied here in T2D, this approach represents a widely applicable high-throughput framework for refining the genetic architecture of complex traits.
nakajima, K.; Sekine, A.
Show abstract
Hypertension is commonly defined as a binary condition despite substantial heterogeneity in diagnosis, treatment, and blood pressure (BP) control. We propose a three-axis state model integrating diagnosis status, treatment intensity, and BP control to better characterize hypertension phenotypes. The framework generates 27 possible states that can be condensed into seven clinically meaningful groups. We applied the model to 5,129,584 Japanese adults using the National Database of Health Insurance Claims and Specific Health Checkups. Hierarchical cluster analysis, sensitivity analysis excluding patients with cardiovascular diseases other than hypertension, and validation against antihypertensive medication use were performed. Overall, 64% of participants were classified as normotensive, whereas 36% belonged to hypertension-related groups, including 11% with unrecognized hypertension and 7% with diagnosed but untreated hypertension. Agreement with data-driven hierarchical cluster analysis was substantial (weighted {kappa}=0.87). The group distribution remained largely unchanged in the sensitivity analysis, supporting the robustness of the proposed classification. Hypertension diagnosis also showed high validity, with a sensitivity of 96.5%, specificity of 91.8%, and substantial agreement with antihypertensive medication use ({kappa}=0.78). This three-axis framework provides a robust and clinically interpretable approach for characterizing hypertension phenotypes, enabling systematic identification of care gaps and supporting research, clinical decision-making, and population health management.
Hsu, C.-Y.; Liu, Q.; Shyr, Y.
Show abstract
As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.
Amato, L. G.; Angiolelli, M.; Demuru, M.; Troisi Lopez, E.; Quarantelli, M.; Granata, C.; Depannemaecker, D.; Jirsa, V.; Bonavita, S.; Mazzoni, A.; Sorrentino, P.
Show abstract
Comprehensive biomarkers of multiple sclerosis (MS) capable of simultaneously diagnosing the condition, capturing symptom severity and predicting treatment efficacy remain elusive. Although several studies have highlighted the pivotal role played by demyelinating lesions in determining MS structural pathology, their relationship with symptom severity is limited. Here, we combined personalized computational brain modeling with magnetoencephalography (MEG) recordings from 17 MS patients and 20 healthy controls (CTR) to derive personalized brain network excitability parameters, which we tested as MS biomarkers. Personalized parameters discriminated between CTR and MS participants with high accuracy, also classifying between progressing and remitting MS patients. Notably, they also predicted MS clinical scales across multiple domains. In all clinical tasks, personalized parameters consistently outperformed standard clinical measures and total lesion loads. Together, these results highlight the potential of personalized brain modelling in deriving integrative MS biomarkers, capable of simultaneously identifying the condition, classifying MS subtypes and predicting symptom severity. d brain modelling in deriving integrative MS biomarkers, capable of simultaneously identifying the condition, classifying between MS subtypes and predicting the severity of symptomatology.
Azizi, L.; Aksoylu, I.; Bueno Alvez, M.; Foucher, J.; Juto, A.; Seitz, C.; Press, R.; Samuelsson, K.; Kläppe, U.; Uhlen, M.; Edfors, F.; Bergström, S.; Fang, F.; Nilsson, P.; Öijerstedt, L.; Manberg, A.; Ingre, C.
Show abstract
Background: Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disease characterized by death of upper and lower motor neurons, usually presented with clinical heterogeneity. Fluid biomarker development remains dominated by neurofilament light chain (NEFL), a marker of neuroaxonal injury. NEFL is however unspecific to ALS and its phenotypes and there is currently a lack of biomarkers that capture ALS heterogeneity such as onset site and ALS-frontotemporal spectrum disorder (ALS-FTSD). Therefore, we investigated whether plasma proteomics could reveal pathway-level signatures that stratify and explain ALS heterogeneity. Methods: We profiled ~5,400 plasma proteins (Olink Explore HT) in 299 patients with ALS and 50 age- and sex comparable healthy controls. We used two complementary analytic frameworks: (i) differential protein abundance analysis to identify altered proteins in ALS and across clinical subgroups, and (ii) weighted gene correlation network analysis (WGCNA) to identify coordinated protein modules and relate them to ALS diagnosis and to ALS-specific clinical traits (site of onset, ALS-FTSD, ALS functional rating scale-revised (ALSFRS-R) score, and plasma NEFL). Results: Differential abundance analysis identified 56 proteins altered in ALS versus controls, of which 40 were increased. WGCNA identified 11 co-expression modules, with ALS samples having the strongest correlation to a protein module (n=51) highly enriched for muscle-related proteins. Out of the 40 proteins that had increased expression levels, 29 overlapped with the muscle-enriched protein module, indicating that muscle related proteins are the dominant circulating proteomic signature in ALS. This signal extended to clinical stratification: spinal-onset patients showed a strong positive association with the muscle-module. Further, differential abundance analysis of spinal- versus bulbar-onset ALS identified changes that mapped predominantly to the same module, supporting a molecular signature of onset phenotype. In contrast, cognitive status (ALS-FTSD) mapped to distinct modules enriched for extracellular matrix/cell-adhesion pathways, consistent with a separable biological axis of disease heterogeneity. Although multiple modules correlated with NEFL, trait-specific signatures were not fully explained by neuroaxonal injury. Notably, the muscle-enriched module increased with higher NEFL and lower ALSFRS-R, supporting its interpretation as a severity-linked, muscle-involvement proxy. Conclusions: Large-scale plasma proteomics reveals that heterogeneity in ALS reflects underlying biological structures. We identified a dominant muscle-associated protein network that distinguished ALS patients from controls and correlated with disease onset phenotype and severity, alongside distinct protein networks linked to ALS-FTSD. By integrating differential protein abundance with network-based analysis, we defined pathway-level biomarker signatures that extend beyond NEFL, enabling biologically informed patient stratification and improved therapeutic monitoring.